Turkish LVCSR: Database Preparation and Language Modeling for an Agglutinative Language
نویسندگان
چکیده
Turkish language is an agglutinative language. It is possible to produce a very high number of words from the same root with suffixes [1]. Language modeling for agglutinative languages needs to be different than modeling of languages like English. Such languages also have inflections but not as many as an agglutinative language. Techniques which can be used for modeling agglutinative languages are presented in this work. Turkish is one of the least studied language for speech recognition. For this reason the first step for Turkish speech recognition is preparing a database. The texts to record the database were selected from television programs and newspaper articles. Selection criterion was to cover various subject and to create a phonetically balanced corpus. Additionally it is important to include as many different word as possible. The Speech Training and Recognition Unified Tool (STRUT)1 has been used for training and testing systems for preliminary recognition experiments.
منابع مشابه
Turkish LVCSR: towards better speech recognition for agglutinative languages
The Turkish language belongs to the Turkic family. All members of this family are close to one another in terms of linguistic structure. Typological similarities are vowel harmony, verb-final word order and agglutinative morphology. This latter property causes a very fast vocabulary growth resulting in a large number of out-of-vocabulary words. In this paper we describe our first experiments in...
متن کاملOn morph-based LVCSR improvements
Efficient large vocabulary continuous speech recognition of morphologically rich languages is a big challenge due to the rapid vocabulary growth. To improve the results various subword units called as morphs are applied as basic language elements. The improvements over the word baseline, however, are changing from negative to error rate halving across languages and tasks. In this paper we make ...
متن کاملOn lexicon creation for turkish LVCSR
In this paper, we address the lexicon design problem in Turkish large vocabulary speech recognition. Although we focus only on Turkish, the methods described here are general enough that they can be considered for other agglutinative languages like Finnish, Korean etc. In an agglutinative language, several words can be created from a single root word using a rich collection of morphological rul...
متن کاملCoalescence Type based Confidence Warping for Agglutinative Language Keyword Spotting
In agglutinative languages like Korean, words are formed by joining l affix morphemes to the stem, which leads to high OOV rate in dictionary building. Hence, subword units are usually used as basic language modeling units in Large-Vocabulary Continuous Speech Recognition (LVCSR) or LVCSR based applications such as keyword spotting. In this work, firstly a new word property called coalescence t...
متن کاملLattice extension and rescoring based approaches for LVCSR of Turkish
In this paper, we present some techniques to solve the problems of Turkish Large Vocabulary Continuous Speech Recognition (LVCSR). Its agglutinative nature makes Turkish a challenging language in terms of speech recognition since it is impossible to include all possible words in the recognition lexicon. Therefore, data-driven sub-word recognition units, in addition to words, are used in a newsp...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2001